Papers with text mining methods
ILCM - A Virtual Research Infrastructure for Large-Scale Qualitative Data (L18-1)
Copied to clipboard
Andreas Niekler, Arnim Bleier, Christian Kahmann, Lisa Posch, Gregor Wiedemann, Kenan Erdogan, Gerhard Heyer, Markus Strohmaier
| Challenge: | iLCM project develops integrated research environment for qualitative data analysis . text mining and text mining tools are extended by "Open Research Computing" |
| Approach: | iLCM project develops integrated research environment for analysis of structured and unstructured data in a "Software as a Service" architecture. |
| Outcome: | iLCM project develops integrated research environment for analysis of structured and unstructured data in a "Software as a Service" architecture. |
Automatic Identification of Research Fields in Scientific Papers (L18-1)
Copied to clipboard
Eric Kergosien, Amin Farvardin, Maguelonne Teisseire, Marie-Noëlle Bessagnet, Joachim Schöpfel, Stéphane Chaudiron, Bernard Jacquemin, Annig Lacayrelle, Mathieu Roche, Christian Sallaberry, Jean Philippe Tonneau
| Challenge: | TERRE-ISTEX project aims to identify scientific research dealing with specific geographical territories areas based on heterogeneous digital content available in scientific papers. |
| Approach: | TERRE-ISTEX project aims to identify scientific research dealing with specific geographical territories areas based on heterogeneous digital content available in scientific papers. |
| Outcome: | The proposed method will help scientists identify geographical territories areas from scientific papers available in digital versions within and outside the ISTEX library. |
Tri-Train: Automatic Pre-Fine Tuning between Pre-Training and Fine-Tuning for SciNER (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Pre-training a language model by self-supervised tasks on huge datasets and fine-tuning with small labelled data are often inadequate for scientific NER tasks. |
| Approach: | They propose to introduce a "pre-fine tuning" step between pre-training and fine-tuning to construct a corpus by selecting sentences from unlabeled documents that are the most relevant with labelled training data. |
| Outcome: | The proposed approach improves on seven benchmarks on the performance of the proposed model on labelled datasets. |